RefConcile - Automated Online Reconciliation of Bibliographic References

نویسندگان

  • Guido Sautter
  • Klemens Böhm
  • David King
چکیده

Comprehensive bibliographies often rely on community contributions. In such settings, de-duplication is mandatory for the bibliography to be useful. Ideally, de-duplication works online, i.e., when adding new references, so the bibliography remains duplicate-free at all times. While de-duplication is well researched, generic approaches do not achieve the result quality required for automated reconciliation. To overcome this problem, we propose a new duplicate detection and reconciliation technique called RefConcile. Aiming specifically at bibliographic references, it uses dedicated blocking and matching techniques tailored to this type of data. Our evaluation based on a large realworld collection of bibliographic references shows that RefConcile scales well, and that it detects and reconciles duplicates highly accurately.

برای دانلود رایگان متن کامل این مقاله و بیش از 32 میلیون مقاله دیگر ابتدا ثبت نام کنید

ثبت نام

اگر عضو سایت هستید لطفا وارد حساب کاربری خود شوید

منابع مشابه

Comparison of Bibliographic Databases in Retrieving Information on Telemedicine

Background & Aims: Some of the main questions which can be of importance for those researchers who intend to perform a systematic review in a field of science are: ‘What databases should I use for my review?’; ‘Do all these databases have the same value?’; and ‘Which sourcesretrieved the highest of relevant references?’. The main aim of this work was the identification of the best database for ...

متن کامل

A review of medication reconciliation issues and experiences with clinical staff and information systems.

Medication reconciliation was developed to reduce medical mistakes and injuries through a process of creating and comparing a current medication list from independent patient information sources, and resolving discrepancies. The structure and clinician assignments of medication reconciliation varies between institutions, but usually includes physicians, nurses and pharmacists. The Joint Commiss...

متن کامل

Automated Resolution of Noisy Bibliographic References

We describe a system used by the NASA Astrophysics Data System to identify bibliographic references obtained from scanned article pages by OCR methods with records in a bibliographic database. We analyze the process generating the noisy references and conclude that the three-step procedure of correcting the OCR results, parsing the corrected string and matching it against the database provides ...

متن کامل

An automated method for the layup of fiberglass fabric

................................................................................................................................... v CHAPTER 1 GENERAL INTRODUCTION ................................................................................ 1 1.1 Background ........................................................................................................................... 1 1.2 Moti...

متن کامل

Automated Labeling Of Biomedical Online Journal Articles

An automated labeling (AL) module has been developed to automate the extraction of bibliographic data (e.g., article title, authors, affiliation, abstract, and others) from online biomedical journals for the National Library of Medicine’s MEDLINE database. The AL module employs string matching, statistics, and fuzzy rule-based algorithms to identify segmented zones in an article’s HTML pages a...

متن کامل

ذخیره در منابع من


  با ذخیره ی این منبع در منابع من، دسترسی به آن را برای استفاده های بعدی آسان تر کنید

عنوان ژورنال:

دوره   شماره 

صفحات  -

تاریخ انتشار 2013